Why Even "95% Accurate" Agents Are Usually Wrong
European insurers pay €2.8 billion a day in claims. At 95% per-step accuracy, an 8-step agentic workflow gets one in three wrong.
PwC and Microsoft published Unlocking Tomorrow, a playbook on agentic AI in financial services. It describes autonomous agents orchestrating claims journeys from first notice of loss to payment initiation, multi-agent platforms handling drawdowns, rollovers and repayments, and straight-through underwriting below defined thresholds.
The governance chapter covers regulations on Responsible AI, DORA, and the EU AI Act, and calls for a control tower to supervise agent behaviour. This edition is about challenges agentic AI poses to financial reporting and how to address those.
The arithmetic nobody runs
Every AI-agent demo comes with an accuracy number. The number is almost always a single-step figure: 85%, 95%, 99%. The number is real. It is also misleading.
Multi-step agentic workflows compound probability at every step. A workflow with eight steps running at 95% per-step accuracy produces end-to-end correctness of roughly 66%. Patronus AI's TRAIL benchmark puts real-world multi-step accuracy even lower: around 11% on realistic trace, implying 73% per-step accuracy results in 1 out of 9 workflows completing successfully end-to-end. The arithmetic illustrates how even small per-step error rates compound across Agentic workflows. CFOs who engage with the governance architecture now, before deployments reach financial reporting processes, are in a different position from those who engage after.
Why traditional controls don't reach
Deloitte's Q4 2025 CFO Signals survey of 200 finance chiefs at companies with at least $1 billion in revenue finds that 87% of CFOs believe AI will be extremely or very important to their finance department's operations in 2026, and 54% cite integrating AI agents as a top transformation priority. At the same time Deloitte reported only 1 in 5 organisations have a mature governance model in place for autonomous AI agents. The question is how control architecture in both upstream and finance processes will impact financial reporting.
Compounding error rates change the risk calculus in three ways that traditional IT General Controls do not cover.
First, the reliability chain breaks. Traditional validation depends on the principle that the same input produces the same output. Reconcile outputs to inputs, recalculate results, verify against source documents. Probabilistic systems break that principle by design. The control framework has to shift from "verify the logic once and rely on it" to "monitor the distribution of outputs and detect drift before it becomes material."
Second, errors cascade rather than isolate. A deterministic system fails in one place. A probabilistic agent operating in a chain can produce outputs that remain individually plausible but collectively wrong, with the error introduced at step three propagating through steps four through eight without any single step looking anomalous. Monitoring has to happen continuously throughout the process chain, not just at the point where outputs reach the ledger.
Third, most of these agents sit upstream of finance. The claims agent, the underwriting agent, the dynamic pricing agent. Finance does not build them or own them. Finance may not know when they are modified. But finance owns the misstatement risk when they fail. Zillow took a $304 million write-down in 2021 on an upstream valuation model that finance did not build. The control failure was not in the algorithm. It was in the absence of a control architecture connecting a probabilistic upstream system to the balance sheet it eventually touched.
The architecture that closes the gap
The control problem is structural, and it has a structural answer. The architecture has three components that work together.
Process decomposition before agent selection. Before evaluating any agentic solution, decompose the target process into discrete tasks. As discussed in the previous edition, for each task, determine whether it belongs in deterministic automation, probabilistic AI, or human judgment. The compounding error math applies only to the probabilistic steps. Finance workflows are almost always mixtures. The first governance move is knowing which steps are which, because the control framework differs for each. An invoice extraction step is deterministic. A coverage calculation step that draws on policy terms, claim history, and behavioural signals is probabilistic. Treating them identically in governance is the gap.
Governance intensity matched to agent profile. Not all agents require identical oversight. Classification should run across four dimensions: autonomy level, predictability of outputs, decision authority, and system integration scope. An invoice data extraction agent with constrained autonomy and predictable outputs warrants lighter governance than a treasury cash optimisation agent with high autonomy and multi-system reach. The classification determines what happens next: what approval gates are required, what escalation triggers are defined, and what override protocols allow finance to intervene. These elements belong in the agent's architecture from inception. Retrofitting governance onto deployed agents creates gaps and introduces friction at the worst possible time.
Dual-path monitoring that detects drift before it reaches the ledger. The PwC/Microsoft control tower concept is positioned as infrastructure for recording and monitoring agentic workflows. That framing is useful but incomplete for ICFR purposes. The monitoring architecture needs two distinct paths.
Hot path monitoring captures operational signals in near real-time: retrieval latency, errors, rejected operations, governance violations. It answers the question "is something failing right now?" Warm path monitoring aggregates telemetry over time to detect patterns that don't manifest as discrete failures: response consistency drift, reasoning loops, gradual efficiency decline, increasing exception rates. It answers the question "is performance degrading in a way that will matter at period close?"
The distinction matters for financial reporting because the failure mode most relevant to ICFR is not the dramatic one. It is the slow drift. A model that processes claims correctly 95% of the time in January, 92% in March, and 88% in June may never trigger a hot path alert. It will, eventually, produce a misstatement. Warm path monitoring is what catches that trajectory before the auditor does.
Where the human sits
PwC recommends a human-in-the-loop approach that requires human sign-off on agent actions exceeding defined thresholds based on risk factors, transaction values, or other considerations. That framing describes a threshold trigger. The governance question goes further: for which outputs does a human sit between the agent and a consequential decision, and for which outputs does a human supervise the aggregate rather than each transaction?
Human-in-the-loop is appropriate for high-materiality outputs where each individual decision carries misstatement risk: a large settlement determination, a treasury transaction above a defined threshold, a pricing adjustment to a material account. Human-on-the-loop is appropriate for high-volume outputs where the individual transaction is immaterial but the aggregate pattern is: claims frequency drift, coverage calculation distribution shifts, fraud flag rate changes. The control activity is different in each case. The first requires an approval gate. The second requires a monitoring protocol with defined escalation thresholds.
The question finance needs to answer for each agent deployment is not "is a human involved?" but "at what point does human accountability attach, and what information does that human have when it does?"